New analysis of the asymptotic behavior of the Lempel-Ziv compression algorithm
نویسندگان
چکیده
We give a new analysis and proof of the Normal limiting distribution of the number of phrases in the 1978 Lempel-Ziv compression algorithm on random sequences built from a memoriless source. This work is a follow-up of our last paper on this subject in 1995. The analysis stands on the asymptotic behavior of a DST obtained by the insertion of random sequences. Our proofs are augmented of new results on moment convergence, moderate and large deviations, redundancy analysis. Key-words: data compression, digital search tree, generating functions, depoissonization, renewal processes in ria -0 04 76 90 2, v er si on 1 27 A pr 2 01 0 Nouvelle analyse du comportement asymptotique de l’algorithme de compression de Lempel-Ziv Résumé : Nous donnons une nouvelle analyse et preuve de la convergence vers la loi normale des performances de l’algorithme de compression de Lempel et Ziv 1978 sur des séquences aléatoires construites à partir d’une source sans mémoire. Ce travail est une continuation de notre papier sur le même sujet de 1995. L’analyse repose sur le comportement asymptotique d’un arbre digital de recherche obtenu par l’insertion de séquences aléatoires. Nos preuves sont augmentées de nouveaux résultats sur la convergence des moments, les moyennes et grandes déviations et l’analyse de la redondance. Mots-clés : compression de données, arbre digital de recherche, fonctions génératrices, dépoissonisation, processus à renouvellement in ria -0 04 76 90 2, v er si on 1 27 A pr 2 01 0 New analysis of Lempel-Ziv algorithm 3 1 Ziv Lempel compression algorithm We investigate the number of phrases generated by the compression algorithm Lempel-Ziv 1978 [?]. The algorithm consists in arranging the text to be compressed in consecutive phrases such that the next phrase is the unique prefix of the remaining text that has a previous phrase as largest prefix. In other words the next phrase can be identified by its prefix in the previously listed phrases plus one symbol: namely one pointer and a symbol. It is convenient to see the phrases organized in a digital search tree. Imagine a root. The first phrase is the first symbol, say a pending on the root. Depending if the next symbol is again a a, then the next phrase will be with two symbols, the last symbol being appended to the node with a. The new phrase can be read from the root to the new symbol. If the next symbol was different of a, say b, then the new phrase is simply b (the largest phrase prefix is therefore empty) and is directly appended to root. Let a text w written on alphabet A, and T (w) is the digital tree created by the algorithm. To each node in T (w) corresponds a phrase in the parsing algorithm. Let L(w) be the path length of T (w): the sum of all the node distances to the root. We should have L(w) = |w|. To make this rigorous we have to assume that the text w ends with an extra symbol that never occurs before in the text. The knowledge of T (w) without the node sequence order is not sufficient to recover the original text w. But if we know the sequence order of node creation in the tree, then we can reconstruct the original text w. The compression code C(w) is in fact a description of T (w), node by node in their order of creation, each node being identified by a pointer to its parent node in the tree and the symbol that label the edge that link the parent to the node. Since this description contains the sequence order of node creation, then the original text can be recovered. Every node points to parents that have been inserted before, therefore pointer for the kth node costs at most dlog2 ke, and the next symbol costs dlog2 |A|e. A negligible economy can be done by fusionning the two informations in a pointer of cost dlog2(k|A|)e. The compressed code has therefore length
منابع مشابه
Source coding, large deviations, and approximate pattern matching
In this review paper, we present a development of parts of rate-distortion theory and pattern-matching algorithms for lossy data compression, centered around a lossy version of the asymptotic equipartition property (AEP). This treatment closely parallels the corresponding development in lossless compression, a point of view that was advanced in an important paper of Wyner and Ziv in 1989. In th...
متن کاملThe Redundancy and Distribution of the PhraseLengths of the Fixed - Database
The Fixed-Database version of the Lempel Ziv algorithm closely resembles many versions that appear in practice. In this paper, we ascertain several key asymptotic properties of the algorithm as applied to sources with nite memory. First, we determine that for a dictionary of size n, the algorithm achieves a redundancy n = H log log n log n + o(log log n log n), where H is the entropy of the pro...
متن کاملPushdown Compression
The pressing need for efficient compression schemes for XML documents has recently been focused on stack computation [6, 9], and in particular calls for a formulation of information-lossless stack or pushdown compressors that allows a formal analysis of their performance and a more ambitious use of the stack in XML compression, where so far it is mainly connected to parsing mechanisms. In this ...
متن کاملA New Approach to Detect Congestive Heart Failure Using Symbolic Dynamics Analysis of Electrocardiogram Signal
The aim of this study is to show that the measures derived from Electrocardiogram (ECG) signals many a time perform better than the same measures obtained from heart rate (HR) signals. A comparison was made to investigate how far the nonlinear symbolic dynamics approach helps to characterize the nonlinear properties of ECG signals and HR signals, and thereby discriminate between normal and cong...
متن کاملA New Approach to Detect Congestive Heart Failure Using Symbolic Dynamics Analysis of Electrocardiogram Signal
The aim of this study is to show that the measures derived from Electrocardiogram (ECG) signals many a time perform better than the same measures obtained from heart rate (HR) signals. A comparison was made to investigate how far the nonlinear symbolic dynamics approach helps to characterize the nonlinear properties of ECG signals and HR signals, and thereby discriminate between normal and cong...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
عنوان ژورنال:
دوره شماره
صفحات -
تاریخ انتشار 2010